Weakly Supervised Clustering: Learning Fine-Grained Signals from Coarse Labels

نویسندگان

  • Stefan Wager
  • Alexander Blocker
  • Niall Cardin
چکیده

Consider a classification problem where we do not have access to labels for individual training examples, but only have average labels over subpopulations. We give practical examples of this setup, and show how these classification tasks can usefully be analyzed as weakly supervised clustering problems. We propose three approaches to solving the weakly supervised clustering problem, including a latent variables model that performs well in our experiments. We illustrate our methods on an industry dataset that was the original motivation for this research.

برای دانلود رایگان متن کامل این مقاله و بیش از 32 میلیون مقاله دیگر ابتدا ثبت نام کنید

ثبت نام

اگر عضو سایت هستید لطفا وارد حساب کاربری خود شوید

منابع مشابه

A Brief Introduction to Weakly Supervised Learning

Supervised learning techniques construct predictive models by learning from a large number of training examples, where each training example has a label indicating its ground-truth output. Though current techniques have achieved great success, it is noteworthy that in many tasks it is difficult to get strong supervision information like fully ground-truth labels due to the high cost of data lab...

متن کامل

A Weakly Supervised Model for Sentence-Level Semantic Orientation Analysis with Multiple Experts

We propose the weakly supervised MultiExperts Model (MEM) for analyzing the semantic orientation of opinions expressed in natural language reviews. In contrast to most prior work, MEM predicts both opinion polarity and opinion strength at the level of individual sentences; such fine-grained analysis helps to understand better why users like or dislike the entity under review. A key challenge in...

متن کامل

Semi-supervised latent variable models for sentence-level sentiment analysis

We derive two variants of a semi-supervised model for fine-grained sentiment analysis. Both models leverage abundant natural supervision in the form of review ratings, as well as a small amount of manually crafted sentence labels, to learn sentence-level sentiment classifiers. The proposed model is a fusion of a fully supervised structured conditional model and its partially supervised counterp...

متن کامل

NUS-PT: Exploiting Parallel Texts for Word Sense Disambiguation in the English All-Words Tasks

We participated in the SemEval-2007 coarse-grained English all-words task and fine-grained English all-words task. We used a supervised learning approach with SVM as the learning algorithm. The knowledge sources used include local collocations, parts-of-speech, and surrounding words. We gathered training examples from English-Chinese parallel corpora, SEMCOR, and DSO corpus. While the fine-grai...

متن کامل

Non-expert Labels Improve Fine-Grained Object Recognition

Fine-grained object recognition typically requires domain experts to provide class label annotations, making these labels difficult to collect when experts are rare. Often, fine-grained classes can be organized into a visual class taxonomy according to shared visual features. In this case, non-experts can easily distinguish between coarse classes, allowing them to annotate examples at a coarse ...

متن کامل

ذخیره در منابع من


  با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید

عنوان ژورنال:

دوره   شماره 

صفحات  -

تاریخ انتشار 2013